Papers by Venkata S Govindarajan
Measuring Lexical Diversity of Synthetic Data Generated through Fine-Grained Persona Prompting (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Fine-grained personas have been used for generating ‘diverse’ synthetic data for pre-training and supervised fine-tuning of Large Language Models (LLMs). |
| Approach: | They measure the diversity of persona-driven synthetically generated prompts and responses with a suite of lexical diversity and redundancy metrics. |
| Outcome: | The proposed model is based on human-written prompts and responses, but human-generated prompts are significantly less diverse than human-created ones. |
Dark & Stormy: Modeling Humor in Sentences from the Bulwer-Lytton Fiction Contest (2026.acl-short)
Copied to clipboard
| Challenge: | a corpus of "bad" humor sentences from the Bulwer-Lytton Fiction Contest 1 is presented . standard humor detection models perform poorly on corpus, and these sentences combine features common in existing humor datasets with metaphor, metafiction and simile. |
| Approach: | They propose to analyze a corpus of "bad" humor sentences from the Bulwer-Lytton Fiction Contest . they use literary devices to synthesize contest-style sentences that imitate the form but exaggerate the effect . |
| Outcome: | The proposed corpus of sentences from the Bulwer-Lytton Fiction Contest 1 is analyzed . it shows that the sentences combine features common in existing humor datasets with metaphor, metafiction and simile. |